Back

European Journal of Human Genetics

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match European Journal of Human Genetics's content profile, based on 58 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
A Randomized Non-Inferiority Trial of an eHealth Delivery Alternative for Cancer Genetic Testing for Hereditary Cancer (eREACH2)

Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361920 medRxiv
Top 0.1%
6.9%
Show abstract

Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.

2
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.3%
3.2%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

3
A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders

Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.

2026-08-31 health informatics 10.64898/2026.08.28.26361655 medRxiv
Top 0.6%
1.7%
Show abstract

Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.

4
Hybrid risk scores integrating polygenic and clinical variables for endometriosis prediction

Goroshchuk, O.; Koller, D.

2026-09-03 epidemiology 10.64898/2026.08.31.26361798 medRxiv
Top 0.8%
1.1%
Show abstract

Background: Endometriosis affects approximately 10% of reproductive-age women and is associated with substantial diagnostic delay and heterogeneous symptom presentation. Prior machine-learning prediction models have relied on comorbidity data alone or on small candidate-variant genetic scores, with inconsistent or incompletely reported performance. No study has combined a well-powered, multi-ancestry polygenic risk score (PRS) with environmental, reproductive, and symptom data in a single hybrid model. We developed and evaluated hybrid risk-prediction models integrating a genome-wide, multi-ancestry PRS with clinical and symptom data for endometriosis in the US-based All of Us Research Program. Methods: Among 69,376 participants (15,382 endometriosis cases, 53,994 controls) across six genetically inferred ancestry groups, we computed individual-level PRS values using PRS-CS weights derived from an independent, multi-ancestry GWAS. Five nested logistic regression, random forest, and XGBoost models progressively added age, ancestry, and within-ancestry genetic principal components (Model 1), environmental and reproductive factors (Model 2), symptom and comorbidity indicators (Model 3), all covariates combined (Model 4), and PRS x environment interactions (Model 5). Performance was assessed by AUROC in a held-out test set and 5-fold cross-validation, with class-weighted, Youden-optimized thresholds used for sensitivity, specificity, and predictive values; permutation importance identified top contributors. Pairwise AUROC differences were tested with a Holm-corrected DeLong-type test. Results: Discrimination improved from AUROC 0.63 (PRS, age, ancestry, principal components) to 0.72 for the full model, driven mainly by symptom and comorbidity data. XGBoost consistently outperformed logistic regression and random forest. The PRS ranked among the top individual predictors by permutation importance in nearly every model, alongside age, while genetic and demographic information alone gave only modest discrimination, and PRS x environment interactions did not improve on environmental factors alone. Threshold optimization yielded balanced sensitivity and specificity (~0.67/0.65) versus near-zero sensitivity at a default threshold. Conclusions: Combining the PRS with symptom and comorbidity data gave the best discrimination compared to solely a well-powered, multi-ancestry PRS as a predictor of endometriosis. This study clarifies both the promise and current limits of hybrid genetic-clinical prediction for endometriosis and points to symptom-based phenotyping, molecular subtyping, and external validation as priorities.

5
Constitutive PDGFRb activation drives connective tissue overgrowth through STAT5-IGF1 signaling

Kwon, H. R.; Rackley, A.; Olson, L. E.

2026-08-29 genetics 10.64898/2026.08.27.747555 medRxiv
Top 0.8%
1.1%
Show abstract

Autosomal dominant gain-of-function mutations in platelet-derived growth factor receptor beta (PDGFRb) cause overgrowth of the skeleton and other connective tissue in Kosaki overgrowth syndrome. However, the target cell type and signaling pathways underlying PDGFRb-driven overgrowth are unknown. Normal postnatal growth is controlled by pituitary-secreted growth hormone (GH), which activates the STAT5 transcriptional factor to upregulate insulin-like growth factor 1 (IGF1). To investigate the role of the GH-STAT5-IGF1 pathway in PDGFRb-related overgrowth, we generated mice with a PDGFRb gain-of-function mutation in skeletal and fibroblast lineages, which resulted in STAT5 activation and gigantism. Conditional deletion of Stat5ab in connective tissue lineages rescued skeletal overgrowth and keloid-like fibrosis in the skin. Conditional deletion of GH receptor (Ghr) did not rescue overgrowth, indicating the physiological activator of STAT5 is not required for overgrowth. However, deletion of Igf1, the STAT5 target gene, and its receptor, Igf1r, in connective tissue, rescued the overgrowth phenotype. These findings demonstrate a GHR-independent STAT5-IGF1 signaling pathway in mutant connective tissue cells, which mediates PDGFRb-driven overgrowth in mice and potentially in humans with similar PDGFRB mutations.

6
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 1%
0.9%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

7
A Curated Pharmacogenomic Allele Catalog for Sub-Saharan African Populations

SULAIMAN, M. A.; Oyeyemi, B. F.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361354 medRxiv
Top 1%
0.6%
Show abstract

Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.

8
Projected Population-Level Impact of Digital Return of Results for Cardiovascular-Kidney-Metabolic Screening at US Blood Donation Centers: A Monte Carlo Simulation Study

Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.

2026-09-03 public and global health 10.64898/2026.09.01.26360806 medRxiv
Top 1%
0.6%
Show abstract

Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.

9
Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology

Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.

2026-09-02 hematology 10.64898/2026.09.01.26361881 medRxiv
Top 1%
0.6%
Show abstract

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.

10
Assessing the association between contraceptive agency and preference-aligned fertility management among Ugandan women: A 12-month prospective cohort study

Birabwa, C.; Wasswa, R.; Amongin, D.; Rakesh, G.; Beth, P.; Sneha, C.; Gomez, R.; Atuyambe, L.; Liu, J.; Waiswa, P.; Holt, K.

2026-08-31 sexual and reproductive health 10.64898/2026.08.27.26361582 medRxiv
Top 1%
0.5%
Show abstract

Background There has been a proliferation of new person-centered and human rights-based contraception measures in recent years, though their application in research remains limited. Improved measures offer an opportunity to examine how contraceptive decision-making agency relates to individuals ability to act in line with their contraceptive preferences. We sought to assess the association between contraceptive agency and subsequent Preference-aligned Fertility Management (PFM) over 12 months in a cohort of women in rural Uganda. Methods We analyzed data from a prospective cohort study conducted in five largely rural Ugandan districts from 2022 to 2024. Data were collected at baseline, 6 and 12 months from a convenience sample of women who were new users of contraception or not using contraception. We used mixed-effects logistic regression models to examine the association between baseline Agency in Contraceptive Decisions Scale overall and subscale scores and future PFM Index scores at 6 and 12 months, assessing whether associations varied over time using interaction terms for follow-up time point. We used interactions between agency scores and follow-up visit to assess whether associations differed between the 6- and 12-month visits. We assessed effect modification by age group and baseline contraceptive method category using three-way interaction terms and predicted probabilities. Results The analytic sample comprised 2,227 women. The percentage of women practicing PFM increased from 85.7% at baseline to 93.3% at 12 months. A one-unit increase in Agency in Contraceptive Decisions Scale score was associated with higher odds of subsequent PFM (aOR: 1.68, 95% CI: 1.10-2.54). Subscales 3 (knowledge aligned with preferences) and 4 (control over use or non-use) of the Agency in Contraceptive Decisions Scale were significantly associated with future PFM (aOR: 1.31, 95% CI: 1.04-1.66 and aOR: 1.27, 95% CI: 1.06-1.51, respectively). The association between overall contraceptive agency and PFM did not differ between the 6- and 12-month visits. Three-way interaction tests suggested that the associations between the overall Agency in Contraceptive Decisions Scale score and the PFM outcomes varied jointly by age group and baseline contraceptive method category: overall PFM Index (p<0.001), PFM1 (p=0.011), and PFM2 (p<0.001). Conclusion Our findings suggest that higher levels of contraceptive agency may help women act in line with their contraceptive preferences. Increasing womens knowledge and control over contraceptive use may be particularly essential for preferred contraceptive use. The findings also suggest that the association between contraceptive agency and PFM may vary by womens age group and the method of choice, though further exploration is necessary to examine this influence.

11
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 1%
0.5%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [&ge;] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

12
Young people with obesity and rare disease - genotypes, phenotypes and healthcare use

Chia, C.; Baker, K.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361359 medRxiv
Top 1%
0.5%
Show abstract

Obesity is a significant public health concern. Early-onset obesity in the context of rare disease can reflect genetically-mediated pathology or elevated susceptibility through indirect mechanisms. Mapping the diverse characteristics and needs of young people with obesity in the rare disease population is a first step toward mechanistic and translational research. We carried out a retrospective comparative analysis of demographic, genotypic, phenotypic and health service utilisation data for young people with obesity (cases: n=500) and without obesity (controls: n=11,444) from the UK 100,000 Genomes Project rare disease cohort. Cases and controls were recruited prior to genomic diagnosis, across clinical disorder categories. We observed significant association between socioeconomic deprivation and obesity risk. Young people with obesity had significantly higher utilisations of acute care and mental health services, indicating an overall higher health burden. A curated panel of 519 candidate obesity-associated genes demonstrated aggregate association with obesity, although no single gene reached significance. Phenotypic comparison between cases and controls highlighted increased multi-organ and neurological system involvement, highlighting the overlap between neurodevelopmental and obesity risks. Within the case group, we conducted cluster analysis to identify early-onset obesity groups with different phenotypic profiles, potentially arising from different causal pathways - this identified six obesity subgroups of interest, with differing involvement of neurodevelopmental and other systems. Our study confirms that obesity co-occurs with a wide range of factors within the rare disease population, and is associated with significant physical and mental health needs, requiring holistic lifelong care.

13
Rare and Common Germline and Somatic Variants Shape Immune Cytopenia Risk and Enable Risk Stratification

Faria, S. D. S.; Bineau, J.; Moisan, R.; Legault, M.-A.; Lecluze, E.; Pincez, T.

2026-08-31 hematology 10.64898/2026.08.26.26361484 medRxiv
Top 2%
0.5%
Show abstract

The genetic risk factors of immune cytopenias are unclear. Immune cytopenias have been reported in various genetic contexts: 1) inherited error of immunity genes, mainly due to rare germline variants, 2) systemic lupus erythematosus, associated with common germline variants, 3) hematological malignancies, and 4) clonal hematopoiesis, the latter two due to somatic variants. However, the respective contribution and interaction of these variants remain to be investigated. Here, we used two large biobanks with whole genome sequencing data to systematically investigate the genetic contribution to immune cytopenia. We found that the four types of genetic variants independently contribute to immune cytopenia risk. We notably found that carriers of variants in some autosomal recessive genes of inherited error of immunity had an increased risk of immune cytopenia. Additionally, common variant-mediated risk of systemic lupus erythematosus also increased the risk of immune cytopenia. Overall, a third to a half of patients with immune cytopenia carried at least one of the four genetic risk variants investigated. Combining the four variants allowed stratifying the risk of immune cytopenia in both general and high-risk population. In general population, the 10-year incidence of immune cytopenia in the lowest and highest risk groups was 0.08% and 1.5%, respectively. In sum, this work identified that different genetic risk factors can lead to immune cytopenia. A large proportion of individuals with immune cytopenia carried an underlying genetic risk factor. Finally, combining these genetic risk factors enabled risk stratification.

14
Relation of Self-Reported Race and Genetic Ancestry to Hypertension Prevalence Among Hispanics/Latinos: The Hispanic Community Health Study/Study of Latinos

Montanez-Valverde, R. A.; Kim, V.; Duran-Luciano, P.; Yuan, Y.; Sofer, T.; Kaplan, R. C.; Gallo, L. C.; Talavera, G. A.; Perreira, K. M.; Daviglus, M. L.; Rosas, S. E.; Llabre, M. M.; Elfassy, T.; Li, X.; Isasi, C. R.; Rodriguez, C. J.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361995 medRxiv
Top 2%
0.3%
Show abstract

Background. The imprecision of current metrics to capture the complex genetic admixture and racial identity among Hispanic/Latino individuals in the United States [US] is a concern. We examined the relationship of self-reported race and genetic ancestry with hypertension [HTN] among Hispanics/Latinos. Methods. Cross-sectional study of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), including 10,586 Hispanic/Latino unrelated adults. Genetic ancestry: West African [AA], Amerindian [AI], and European [EA]. Self-reported race: White, Black, Native American, or Multiple/Missing (More than one race or Unknown/Not reported/Refused). HTN: systolic (SBP) [&ge;]130 mmHg, diastolic blood pressure (DBP) [&ge;]80 mmHg, and/or use of HTN medications. Age- and sex adjusted models were used. Results. Self-reported race was White (38{middle dot}6%), Black (3{middle dot}6%), Native American (4{middle dot}1%), and Multiple/Missing (53{middle dot}7%), with Unknown/Not reported/Refused representing 32{middle dot}7%. Black and White Hispanics/Latinos had the greatest AA (55{middle dot}7%) and EA (69{middle dot}3%) ancestries, respectively. Each 10% AA increase was associated with OR 1{middle dot}15, SBP beta +0{middle dot}9 mmHg, and DBP beta +0{middle dot}7 mmHg. Conversely, each 10% AI increase was associated with OR 0{middle dot}83, SBP beta -0{middle dot}4 mmHg, and DBP beta -0{middle dot}6 mmHg. HTN prevalence was highest among those with Black race or in the highest AA quantile (45{middle dot}6% and 48{middle dot}0%, respectively), and lowest among those with Native American race or in the highest AI quantile (37{middle dot}6% and 26{middle dot}7%, respectively). Conclusion. One-third of Hispanics/Latinos did not self-report race. Black or White self-reporting race did somewhat relate to AA or EA ancestry, respectively. HTN profiles were related to self-reported race and genetic ancestry in this admixed population.

15
Machine learning analysis of Autism phenotype data supports a four-dimensional continuum with three overlapping subtypes

Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.27.26361561 medRxiv
Top 2%
0.3%
Show abstract

Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.

16
Socioeconomic position, adverse childhood experiences, and menstrual symptoms in two generations of a prospective UK cohort.

Sawyer, G.; Farooq, B.; Birnie, K.; Fraser, A.; Lawlor, D. A.; Sharp, G. C.; Howe, L. D.

2026-08-31 epidemiology 10.64898/2026.08.27.26361513 medRxiv
Top 2%
0.3%
Show abstract

Background: Inequalities exist for many health outcomes, but there is limited evidence regarding menstrual symptoms despite their importance for health and wellbeing. We aimed to investigate inequalities in menstrual symptoms according to socioeconomic position and childhood adversity. Methods: In two generations (G0 mothers and G1 offspring) from the Avon Longitudinal Study of Parents and Children (ALSPAC), a UK prospective cohort study, we examined associations of multiple indicators of socioeconomic position (SEP) and adverse childhood experiences (ACEs) with menstrual symptoms (pain, abnormal uterine bleeding, and premenstrual syndrome (PMS) measured 3-8-years post-birth in G0 and 17-21-years-old in G1), using multivariable logistic regression. Samples ranged from 4,828 to 9,335 G0 participants and 1,288 to 2,757 G1 participants depending on the exposure-outcome association. Missing data were addressed using multiple imputation and inverse probability weighting. Results: Financial difficulties were associated with greater odds of menstrual pain (G1 OR 1.41; 95% CI 1.07, 1.86: G0 OR 1.55; 95% CI 1.36, 1.76) and irregular cycles (G1 OR 1.60; 95% CI 1.12, 2.29: G0 OR 1.48; 95% CI 1.27, 1.72) in both generations, as well as with short/long cycle lengths in G0 only. Lower education and manual social class were also associated with these three menstrual symptoms in at least one generation. Conversely, higher SEP was associated with PMS in both generations. Higher cumulative ACEs were consistently associated with menstrual pain (4+ compared to none: G1 OR 2.15; 95% CI 1.48, 3.11: G0 OR 1.52; 95% CI 1.29, 1.80) and irregular cycles (G1 OR 1.92; 95% CI 1.20, 3.09: G0 OR 1.54; 95% CI 1.26, 1.87) but not cycle length. Lower parental education, financial difficulties, and cumulative ACEs were associated with heavy bleeding in G1 offspring only, whereas financial difficulties, own manual social class, and cumulative ACEs were associated with prolonged bleeding in G0 mothers only. Higher cumulative ACEs were also associated with PMS in G1 offspring only. Conclusions: We found evidence of inequalities according to socioeconomic disadvantage and childhood adversity for multiple menstrual symptoms, although some associations were only observed in one generation. Findings suggest that menstrual symptoms are disproportionately experienced by socially and socioeconomically disadvantaged women.

17
Plasma and follicular fluid concentrations of carotenoids, tocopherols and retinol in a French population of women undergoing in vitro fertilization: a monocentric non-interventional study

Ndiaye, A.; Thiebaut, A. C. M.; Borel, P.; Sabran, C.; Elis, S.; Guerif, F.; Maillard, V.

2026-09-01 sexual and reproductive health 10.64898/2026.08.28.26360803 medRxiv
Top 2%
0.3%
Show abstract

The distribution of fat-soluble compounds (including antioxidants) in follicular fluid (FF) remains sparsely documented in relation to in vitro fertilization (IVF) outcomes and existing studies have reported diverging associations. This study aimed to describe plasma and FF concentrations of fat-soluble micronutrients in women undergoing IVF and to analyze their adjusted associations with ovarian function, embryo development and pregnancy outcomes. In 2021-2022, plasma and FF samples were collected from 82 women (first IVF cycle) at oocyte puncture, along with lifestyle data covering the three preceding months. Eleven compounds (two tocopherols, three xanthophylls, five carotenes and retinol) were quantified. All compounds were detected in both compartments (lowest in FF) except phytoene, undetectable in FF. Plasma and FF -tocopherol concentrations were positively associated with plasma estradiol levels before oocyte puncture (both p<0.01) while FF -carotene and lycopene were inversely associated with plasma progesterone concentrations (p=0.01 and 0.02, respectively). Plasma phytofluene and phytoene were positively associated with mature oocyte rate (p=0.03 and p=0.01, respectively), while FF retinol was negatively associated (p=0.03). Carotenes, tocopherols and retinol were inversely associated with later IVF outcomes: fertilization rate (p<0.001 for plasma g-tocopherol, 0.02 for FF retinol), top-quality embryo (p=0.02 for plasma phytofluene), biochemical pregnancy at day 7 post-embryo transfer (p=0.05 for plasma -tocopherol, 0.02 for plasma -carotene), clinical pregnancy (p=0.03 for plasma -tocopherol, 0.01 for plasma phytoene) and live birth (p=0.04 for plasma -tocopherol, 0.02 for plasma phytoene). Plasma and FF g-tocopherol were positively associated with embryo fragmentation (both p<0.05). Finally, among xanthophylls, only plasma {beta}-cryptoxanthin was positively associated with plasma progesterone concentrations (p=0.02). Our findings of heterogeneous associations between tocopherols, carotenes, retinol and IVF outcomes across the stages of IVF suggest a beneficial effect limited to early outcomes and support a complex and context-dependent role of these compounds in female reproduction. This manuscript has been submitted to PlosOne on August 19, 2026.

18
Clinical deep sequencing to diagnose pathogenic mosaic variants in malformations of cortical development and epilepsy

Stone, K.; Prinzing, G.; Lai, A.; Smith, L.; Sheidley, B. R.; Corliss, M. M.; Bowling, K.; Cao, Y.; Wiltrout, K.; Stone, S. S. D.; Lidov, H.; Yang, E.; Poduri, A.; D'Gama, A. M.

2026-09-03 neurology 10.64898/2026.09.01.26361943 medRxiv
Top 2%
0.3%
Show abstract

Background and Objectives: Deep sequencing of brain tissue in the research setting has established that mosaic variants are a major cause of malformations of cortical development (MCDs) and epilepsy. However, genetic testing in the clinical setting primarily detects germline variants using clinically accessible samples. We aimed to determine the diagnostic yield and clinical utility of deep sequencing in the clinical setting to identify pathogenic mosaic variants for this population. Methods: We performed a retrospective cohort analysis of individuals at Boston Children's Hospital with MCDs with or without epilepsy who received clinical deep sequencing between September 2017 and February 2026. Demographic, clinical, and genetic testing data were abstracted from the medical record. For individuals without systemic features, we classified brain tissue as an affected tissue sample. For individuals with systemic features, we classified brain or relevant non-brain tissue as affected. The primary outcome was the diagnostic yield of clinical deep sequencing performed using affected vs unaffected tissue samples. The secondary outcome was the clinical utility of genetic diagnoses. Results: Our cohort included 37 individuals (19/37 (51%) female, 18/37 (49%) male) with MCDs, of whom 35/37 (95%) had epilepsy (25 with brain tissue samples available from epilepsy surgery) and 8/37 (22%) had systemic features. Most (35/37 (95%)) had dysplasia phenotypes on MRI and 12/27 (44%) with pathology available had Focal Cortical Dysplasia Type I or II. The diagnostic yield was 53% (17/32; 16 mosaic and 1 germline variant) when clinical deep sequencing was performed using an affected tissue sample vs 0% (0/6) using an unaffected tissue sample (p=0.016). Of the diagnosed cases, 13/17 (76%) had testing performed on brain tissue (1 with systemic features) and 4/17 (24%) on non-brain tissue (3 buccal and 1 duodenal tissue, all with systemic features). All but one diagnosis involved the mTOR pathway. All diagnoses had clinical utility. Discussion: Clinical deep sequencing, when performed using an affected tissue sample, has high diagnostic yield and clinical utility for individuals with MCDs, especially dysplasia phenotypes, and epilepsy. Our findings support implementation of clinical deep sequencing for this population, especially as the genetic diagnoses have implications for emerging precision therapies.

19
Towards Electronic Health Records-Based Paediatric Growth References: Results from the SwissPedGrowth Project

Leuenberger, L. M.; Shoman, Y.; Romero, F.; Sasaki, M.; Deligianni, X.; Goebel, N.; Mozun, R.; Bielicki, J. A.; Burckhardt, M.-A.; Saner, C.; Schwitzgebel, V.; Hauschild, M.; Righini Grunder, F.; Mueller, P.; Schlapbach, L. J.; Jenni, O.; Spycher, B. D.; Kuehni, C. E.; Belle, F. N.; SwissPedHealth consotrium,

2026-09-02 pediatrics 10.64898/2026.08.28.26361619 medRxiv
Top 2%
0.3%
Show abstract

BACKGROUND: We used anthropometric data from electronic health records (EHRs) of Swiss childrens hospitals to evaluate growth references and estimate centile curves. METHODS: We received EHRs extracted from seven Swiss childrens hospitals and analysed two samples: all children with a height, weight, body mass index (BMI), or head circumference recording, and a subsample restricted to children without diseases potentially affecting growth, weighted to represent the general population. We calculated mean z-scores based on the World Health Organization growth references adopted for Switzerland in 2011 (CH-WHO 2011) and current Swiss growth references (Swiss 2026). We estimated sex-specific centile curves in the subsample using generalised additive models for location, scale, and shape. RESULTS: We included 213,868 children with height, 448,002 with weight, 209,244 with BMI, and 67,397 with head circumference recordings. Mean z-scores in the all children sample were (CH-WHO 2011; Swiss 2026): height (0.10; -0.19), weight (0.16; -0.09), BMI (0.04; -0.07), head circumference (-0.28, -0.28); and in the subsample: height (0.34; 0.00), weight (0.27; 0.01), BMI (0.18; 0.05), and head circumference (0.04; 0.01). The 50th height, weight, BMI, and head circumference centiles of girls and boys in the subsample closely followed those of Swiss 2026, with slightly wider 3rd and 97th centiles in infancy and adolescence. CONCLUSION: Height, weight, BMI, and head circumference centiles aligned well with the Swiss 2026 growth references in Switzerland, demonstrating that hospital EHRs could contribute to future growth references.

20
A novel framework leveraging non-causal associations reveals shared pathways linking inflammation and cancer risk

Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.30.26361622 medRxiv
Top 2%
0.3%
Show abstract

Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.